Papers with clean and causal interventions
Mechanistic Analysis Of Universality: Numerical Comparison Circuits Across Transformer Architectures (2026.acl-srw)
Copied to clipboard
| Challenge: | Mechanistic interpretability seeks to identify internal circuits within transformer language models but it is unclear whether they generalize across model families and scales. |
| Approach: | They propose to identify internal circuits within transformer language models by numerical comparisons. |
| Outcome: | The proposed model implementations are consistent across architecture and scale, the authors show . their results highlight the need for cross model comparisons to claim generalization of internal circuits. |